Back

Molecular Biology and Evolution

Oxford University Press (OUP)

Preprints posted in the last 7 days, ranked by how well they match Molecular Biology and Evolution's content profile, based on 542 papers previously published here. The average preprint has a 0.31% match score for this journal, so anything above that is already an above-average fit.

1
Evolutionary origins of protein novelty across an entire yeast subphylum

Tassios, E.; Pyrgelis, N.; Rinker, D.; Tzermpou, E. M.; Hittinger, C. T.; Rokas, A.; Nikolaou, C.; Vakirlis, N.

2026-09-01 genomics 10.64898/2026.08.29.748005 medRxiv
Top 0.3%
18.4%
Show abstract

Genes encoding novel protein sequences are a ubiquitous feature of genomes. They fuel molecular and cellular evolutionary innovations and frequently contribute to species-specific characteristics. We are now unravelling the processes by which they originate, including de novo from noncoding sequences and through extreme divergence, yet how much and what types of novel proteins evolve through each process is still unclear Does the mechanism of origination shape the structural and functional potential of the resulting proteins? Here, we conducted a broad computational investigation of genetic and protein novelty at the scale of the entire subphylum of Saccharomycotina yeasts. We detected more than 5,000 robust de novo genes across 332 species and compared them to more than 10,000 novel genes resulting from extreme sequence divergence, revealing two distinct modes of evolution of novelty. A remarkable 40% of de novo proteins are predicted to localize to mitochondria compared to only 15% of divergent, with the latter also being substantially longer and more disordered. A detailed analysis of conservatively predicted tertiary structures of novel proteins shows that "invention" of novel folds can happen through both processes but is more likely to occur de novo. We also illustrate cases of evolutionary "re-invention" of existing protein folds from non-coding sequences. Our work deepens our understanding of the origins and importance of novel proteins opening new directions for further structural and functional characterization.

2
Ancestral Sequences Cannot be Accurately Reconstructed via Interpolation in a Variational Autoencoder's Latent Space

Gorstein, E.; Tang, M.; Bruzzone, H.; Solis-Lemus, C.

2026-09-01 evolutionary biology 10.1101/2025.11.19.689264 medRxiv
Top 0.4%
15.2%
Show abstract

Standard methods for ancestral sequence reconstruction (ASR) rely on substitution models for the residues in a biological sequence and assume independent evolution across these sites, ignoring the epistatic interactions that shape molecular evolution. In contrast, deep learning models like variational autoencoders (VAEs) can learn low-dimensional representations ("embeddings") of sequences in a protein family that may implicitly handle these dependencies, raising the possibility of performing more accurate ASR by interpolating between extant sequence embeddings within the VAE's latent space. In this study, we test this hypothesis by developing and evaluating a VAE-based ASR pipeline. Benchmarking this approach against established likelihood-based and parsimony methods using various simulations of protein evolution, including scenarios with and without epistasis, we find that the VAE-based approach is consistently and significantly outperformed by standard methods, even in epistatic regimes where it was hypothesized to have an advantage. We further show that this failure is not due to a lack of phylogenetic structure in the latent space, which does contain evolutionary signal. Rather, the primary limitation is the information loss inherent to the autoencoding process: the VAE's decoder cannot generate sequences with sufficient fidelity for the precise demands of ASR.

3
Learning and forecasting shared evolutionary pathways to multi-drug resistance across global pathogens

Aga, O.; Moyo, S.; Ferno, J.; Manyahi, J.; Kibwana, U.; Löhr, I.; Langeland, N.; Blomberg, B.; Johnston, I.

2026-09-01 evolutionary biology 10.64898/2026.08.30.748110 medRxiv
Top 0.5%
14.7%
Show abstract

Infections with bacteria which have evolved multi-drug resistance (MDR) cause millions of deaths worldwide. Large-scale efforts are gathering genotypic and phenotypic data on MDR bacteria, but methods for learning the structure, diversity, and predictors of evolutionary pathways to MDR have yet to take full advantage of these data. Here, we use evolutionary accumulation modelling (EvAM), an emerging class of machine learning methods with roots in cancer progression, to infer these evolutionary pathways across ESKAPEE pathogens (seven bacterial species that dominate health burdens), using a database of over 635k genotyped phenotypic observations from around the world. We identify global patterns in MDR evolutionary pathways, remarkably shared across multiple ESKAPEE species. Species-specific deviations from these stereotypical pathways are connected with geographical and demographic covariates, facilitating predictions of future MDR evolution. We verify these predictions with several hundred new phenotypes from ESKAPEE samples spanning decades of clinical infections in sub-Saharan Africa, demonstrating the capacity to forecast future MDR evolution from these inferred shared pathways.

4
Secondary Structure Diversity of the Mitochondrial Small-Subunit rRNA in Porifera

Zhou, Y.; Gong, L.; Niu, G.; Shi, H.; Gutell, R.; Li, X.; Wei, M.

2026-08-30 evolutionary biology 10.64898/2026.08.28.747467 medRxiv
Top 0.9%
8.8%
Show abstract

Animal mitochondrial rRNAs are commonly viewed as structurally reduced, yet sponge mt SSU rRNAs range from compact to highly expanded structures. Using nine conserved structural anchors, we compared 216 taxonomically resolved records from four classes and 22 orders, including 16 freshwater Spongillida and 200 marine sponges. Twelve homologous hypervariable substructures were coded as structural types, and their ordered combinations as composite types. We identified 38 structural types and 62 composite types across molecules ranging from 853 to 2,019 nt. Hexactinellida and freshwater Spongillida were each uniform for a distinct composite type but differed markedly in overall structure: hexactinellid mt SSU rRNAs were compact, whereas those of Spongillida were long and contained four to five candidate insertion regions. These results show that a conserved scaffold can accommodate extensive lineage-associated structural variation and provide a practical framework for comparing highly divergent mitochondrial rRNAs.

5
Rclade: automated taxonomic collapsing and geological-timescale annotation of time-calibrated phylogenetic trees in R

Zeng, Z.; Wang, Y.

2026-09-01 bioinformatics 10.64898/2026.08.27.747462 medRxiv
Top 1%
6.1%
Show abstract

Background: Reproducible taxonomic collapsing and geological-timescale annotation of time-calibrated phylogenetic trees in R often require coordination among several packages and repeated code for label parsing, clade validation, plotting, and export. Workflow-managed analyses additionally benefit from non-interactive configuration, predictable diagnostics, and machine-readable exit status. Results: We present Rclade, an R package that consolidates the multi-package coordination required for taxonomic collapsing into a streamlined, single-function interface. Rclade provides (1) custom ggproto objects (GeomPolygonStraight/GeomSegmentStraight) that bypass coord_munch() interpolation to achieve straight-edge rendering of collapsed triangles in circular layouts; (2) automatic detection and parsing of four taxonomic-label formats (GTDB, Silva, NCBI, embedded) plus user-supplied custom regex, with explicit input-validation contracts and parsing-accuracy evaluation on real and derived test sets; and (3) workflow embeddability through YAML configuration, library-mode APIs, and standard Unix exit codes. Benchmarks on synthetic and real datasets (200-10,000 synthetic tips and real reference trees up to 10,122 tips; 5 replicates at every scale under a unified fully rendered measurement protocol) show that the full-pipeline overhead is modest for interactive use (median {approx}0.87 s in-session rendering and {approx}8.4 s process-level wall-clock at 10,000 tips). Conclusions: Rclade is a convenience layer over the ggtree/deeptime ecosystem that reduces boilerplate while adding targeted technical improvements for circular-layout rendering and format heterogeneity management.

6
Parallel evolution under constraint shapes echinocandin resistance in Candida auris

Cauldron, N. C.; Dort, E. N.; Weeks, G.; Rogers, D.; Cuomo, C. A. A.

2026-09-01 genetics 10.64898/2026.08.30.748140 medRxiv
Top 1%
5.5%
Show abstract

Drug resistance emerges repeatedly in outbreaks of Candida fungal pathogens, but little is known about its origins or persistence. Here, we investigated the evolutionary processes shaping echinocandin resistance in Candida auris, a globally emerging and predominantly clonal fungal pathogen. Genome-wide association across over 600 isolates identified mutations in the {beta}-1,3-glucan synthase gene FKS1 as the most significant driver of resistance to an echinocandin drug. Ancestral reconstruction of this population traced shared resistance mutations among small groups typically consisting of 2-3 closely related isolates, but clusters could include up to 16 isolates. Nearly all resistant clusters consisted of isolates collected in the same year and region, consistent with local transmission. To further examine population-level selection, we measured adaptive signatures in FKS1 and the highly diverged paralog FKS2 across 22,000 genomes. This revealed excess nonsynonymous polymorphisms in FKS1, primarily due to independent, recurrent mutations at resistance hotspots, consistent with parallel evolution and incomplete fixation of adaptive alleles. In FKS2, there is no evidence of hotspots and little support for diversifying selection. Together, these results indicate that resistance mutations emerge under strong genetic constraint, with adaptation restricted to only one FKS homolog and predominantly at mutational hotspots.

7
Contrasting evolutionary trajectories of nitrate assimilation across Brettanomyces bruxellensis lineages

Vigna, A.; Harrouard, J.; Miot-Sertier, C.; Loegler, V.; Marullo, P.; Friedrich, A.; Schacherer, J.; Peltier, E.; Albertin, W.

2026-08-31 microbiology 10.64898/2026.08.31.748220 medRxiv
Top 2%
4.7%
Show abstract

Brettanomyces bruxellensis is a yeast species associated with diverse fermentation environments and characterized by extensive genetic diversity, including diploid, autotriploid, and allotriploid lineages resulting from independent hybridization events. These lineages are associated with distinct ecological niches and provide a framework for studying metabolic trait evolution in complex genomes. Nitrate assimilation is a relatively uncommon trait among yeasts and has been reported in B. bruxellensis, but its distribution and evolutionary history within the species remain poorly understood. Here, we combined phenotypic characterization of 151 strains with genomic analyses of 946 whole-genome sequences to investigate nitrate assimilation. Growth assays revealed that nitrate assimilation is widespread but unevenly distributed across genetic lineages, with some populations largely retaining the trait whereas others have frequently lost it. Genomic analyses identified extensive variation affecting the nitrate assimilation gene cluster composed of YNR1, YNI1, and YNT1. Nitrate assimilation was strongly associated with both gene copy number and predicted gene functionality, with nitrate-assimilating strains generally carrying more functional copies of the cluster. Leveraging the complex genomic architecture of the species, we independently analyzed primary and acquired genomes in allotriploid lineages and uncovered contrasting evolutionary trajectories following hybridization. While nitrate assimilation genes were generally maintained in primary genomes, acquired genomes showed a higher prevalence of gene loss and predicted loss-of-function variants, revealing asymmetric dynamics between subgenomes. Altogether, our results suggest that nitrate assimilation represents an ancestral trait that has been differentially maintained across B. bruxellensis lineages through a combination of copy number variation, gene degeneration, and genome-specific evolutionary dynamics. These findings provide new insights into how genome architecture and polyploid evolution shape the maintenance and loss of metabolic traits in an industrially relevant yeast species.

8
Cross-Kingdom Control: Yeast Prion Protein Modulates Host Physiology in Drosophila

Clark, A. G.; Jiang, J. Y.; Chitale, M. D.; Cosgrove, E.; Van Elgort, A.; Jain, A. M.; Kelso, J. C.; Cui, X.; Yapici, N.; Lin, C.-c.

2026-09-01 evolutionary biology 10.64898/2026.08.26.747210 medRxiv
Top 2%
3.9%
Show abstract

Prions, once mainly studied for their pathogenic roles, are now gaining recognition as adaptive elements in microbial physiology. Over one-third of wild yeast isolates harbor prion proteins, yet their impact on host-microbe interactions remains poorly characterized. Given the ecological dominance of yeasts in the Drosophila mycobiome, we leveraged the Drosophila melanogaster-Saccharomyces cerevisiae system to investigate how the mycobiome-derived prion, [MRPL10+], modulates host physiology. We show that flies exposed to [MRPL10+] yeast exhibit significantly enhanced cold tolerance and increased locomotor activity. This effect persists with heat-killed yeast and diluted culture, suggesting a stable, potent bioactive factor. Using the genetically diverse Drosophila Global Diversity Lines (GDL), we identified natural variation in responsiveness to [MRPL10+] yeast. Genome-wide association and functional RNAi screening revealed a gut-brain signaling axis involving genes critical for digestion, intercellular communication, transcription regulation, and neural transmission. Notably, serotonin and octopamine pathways were essential for [MRPL10+]-induced changes in cold tolerance and locomotion, implicating neuromodulatory circuits in prion-mediated microbial signaling. Our findings establish a mechanistic link between a fungal prion and host metabolic and neural adaptation. This work provides the first genetic dissection of a prion-mediated host-microbe interaction, laying the groundwork for investigating beneficial prions in complex microbial communities and highlighting a new dimension of the mycobiomes influence on animal physiology.

9
PhageTransformer - scalable and accurate host assignments for bacteriophages

Siemers, M.; Lopez, J. L.; Dutilh, B. E.

2026-08-30 bioinformatics 10.64898/2026.08.29.748026 medRxiv
Top 2%
3.1%
Show abstract

Bacteriophages can only be understood through their interactions with bacterial hosts. As environmental sequencing efforts expanded, the number of available phage genome sequences has exploded, yet the vast majority of these sequences lack host information. Predicting the host of a newly observed phage is therefore a key challenge in virology. Several computational tools can predict phage-host relationships from genomic data, but they share notable limitations: (1) the number of different hosts that can be predicted remains relatively restricted; (2) tools tend to assign confident host predictions to non-viral input sequences; and (3) most tools have a trade-off between accuracy and speed. Here we present PhageTransformer (PT), a deep learning model for phage-host prediction that addresses these limitations. We benchmark PT against existing tools on 3,881 independent phage-host pairs from GenBank and public HiC data, and demonstrate that it achieves competitive or superior prediction accuracy at greatly reduced runtime.

10
Two evolutionary histories in one nucleus: genome remodeling and allelic regulation underlying heterosis in hybrid oil palm

Su, X.; Peng, Y.; Yang, X.; Zhang, F.; Xu, Q.; Ma, Z.; Dong, Y.; Zhou, L.; Xue, H.; Cao, X.; Zou, Z.; Wang, Y.; Zhou, Y.; Zeng, X.

2026-08-31 genomics 10.64898/2026.08.27.747553 medRxiv
Top 2%
3.1%
Show abstract

Oil palm (Elaeis) is the primary source of global vegetable oil. Interspecific hybrids of Elaeis exhibit pronounced heterosis by integrating two distinct subgenomes into a single nucleus, effectively combining the high yield of African oil palm (E. guineensis) with the high unsaturated fatty acid content and disease resistance of American oil palm (E. oleifera). However, the genetic basis underlying heterosis is still unclear. Here, we combine phased genome assembly, comparative genomics, evolutionary genomics and haplotype-aware transcriptomics to unravel the genetic architecture of heterosis of hybrid oil palm. We assemble the highly heterozygous F1 genome ('Reyou 40', 3.75% heterozygosity) into a complete 1.73 Gb T2T haplotype (HapG) and a 1.84 Gb near-T2T haplotype (HapO with17 gaps). Despite 91.56% sequence identity, HapG and HapO diverged in LTR-RT occurrence and PAV affected genes, showing complementary biases in lipid metabolism and stress responses, respectively. Evolutionary genomics revealed that ancient WGDs preserved the palm family. Whereas lineage-specific lipid-related gene expansions in oil palm. Six ancient introgressed regions (~64 Mb) in HapG were reshaped by transposable elements and tandem duplication, showing an enrichment of genes related to resistance and lipid metabolism. Transcriptomically, 82.2% of allelic gene pairs maintained balanced expression, accompanied by parental functional complementarity and dosage buffering, revealing a potential regulatory basis for coordinating parental genetic differences in the hybrid genome. These haplotype-resolved genomic resources offer vital targets for understanding heterosis and accelerating oil palm molecular breeding.

11
GNMCADS: Sampling For Protein Conformation Diversity With Gaussian Network Model Guided Condition Annealed Diffusion Sampler

Uzum, A. S.; Haliloglu, T.

2026-09-01 bioinformatics 10.64898/2026.08.28.747885 medRxiv
Top 3%
2.4%
Show abstract

Proteins are dynamic molecules existing in diverse conformational states underlying their biological functions. Although recent approaches have enabled diverse conformational sampling by emulating molecular dynamics simulations, perturbing evolutionary information, or steering internal mechanisms of structure prediction models, predicting conformations resulting from major domain motions or motions that occur over long timescales still remains a challenge. To this end, we introduce GNMCADS, a conformational sampling strategy that enhances the diversity of protein diffusion models by selectively annealing the conditioning signal guided by the intrinsic dynamical organization of the sampled protein. Further, we implement GNMCADS in the diffusion module of AlphaFold3, enabling the generation of diverse protein conformations. When benchmarked across 92 proteins that include 54 class A GPCRs, 15 transporters, and 23 proteins with major domain movements, GNMCADS exhibits improved sampling diversity compared to other current conformational sampling methods.

12
Discovering 25 novel phyla that fill gaps in the eukaryotic tree of life

Tedersoo, L.; Mikryukov, V.; Sildever, S.; Chmolowska, D.; Piwosz, K.; Meyneng, M.; Monjot, A.; del Campo, J.; Lara, E.; Hakimzadeh, A.; Geisen, S.; Panksep, K.; Bahram, M.; Oliverio, A.; Shepherd, R.; Rückert, S.; Lanzen, A.; Hurdeal, V.; Concetta Eliso, M.; Casotti, R.; Hosseynimoghadam, M.; Siano, R.; Chauvet, M.; Prins, V.; Kisand, V.; Anslan, S.; Alkahtani, S.; Nilsson, H.

2026-08-31 microbiology 10.64898/2026.08.28.747736 medRxiv
Top 3%
2.4%
Show abstract

Protists play important roles in food chains and symbioses in soil and aquatic environments, displaying an enormous morphological and functional diversity. While most commonly found protist species are well known to science, our global-scale environmental DNA survey across soil, water, and sediments reveals dozens of novel, phylum-level phylogenetic lineages that remain to be characterized for basic morphology and function. A vast majority of these undescribed taxa occur in marine water and sediments, but some are common in soil. Most of these novel taxa have distinct substrate and habitat preferences and biogeographic patterns. To accord these lineages scientific agency and enable unambiguous scientific communication, we propose formal names for 150 species to phylum-level taxa from 25 deep lineages based on eDNA and rRNA gene long-read sequence information.

13
The evolutionarily conserved C-terminal domain of a domesticated transposase-derived protein regulates its DNA integration ability

Saha, A.; Ghosh, A.; Majumdar, S.

2026-08-31 biochemistry 10.64898/2026.08.31.747927 medRxiv
Top 3%
2.1%
Show abstract

THAP9 is a transposable element-derived gene which encodes a protein that is homologous to the active Drosophila P-element transposase (DmTNP). Both THAP9 and DmTNP possess a C-terminal domain (CTD) which is functionally uncharacterized. Sequence and structural analysis suggest that the THAP9-CTD has a novel fold which is only found in THAP9 homologs. To explore the evolutionary history and characteristics of this novel domain, exhaustive phylogenetic analysis (using MSA, structure prediction, MSTA-based clustering) was performed. THAP9-CTD homologs were more widely distributed throughout the animal kingdom in comparison to DmTNP-CTD homologs which were restricted to arthropods. Moreover, the THAP9-CTD homologs were more conserved, especially among mammals and birds and their average length increased in a class-specific manner. Comparison with the DmTNP-CTD homologs demonstrates that although their respective CTDs may have evolved independently, they both surprisingly share similar secondary structure elements consisting of three conserved helical regions made of hydrophobic residues that are predicted to make up a conserved core. The role of the respective CTDs were further investigated by creating truncation mutants lacking the CTD. Interestingly both THAP9 and DmTNP truncation mutants are still capable of DNA excision and integration suggesting that their respective CTDs are not essential for DNA transposition. Moreover, CTD truncation favours DNA integration in THAP9: this suggests that CTD acquisition during evolution may have led to THAP9 domestication as observed in other transposable element-derived genes like Rag1 and piggybac, which have similar terminal regulatory domains.

14
Structural characterization the LlaI anti-phage defense system reveals insights into the evolution of nucleotide specificity and the organization of DNA binding in McrBC restriction complexes

Bui, A. Q.; Hosford, C. J.; Niu, Y.; Santiago, E.; Moraga, D.; Wagner, M. M.; Chappie, J. S.

2026-09-01 biochemistry 10.64898/2026.08.31.748284 medRxiv
Top 3%
2.1%
Show abstract

Canonical McrBC enzymes are nucleotide-powered, motor-driven endonucleases that bind and cleave modified bacteriophage DNA. Non-canonical McrBC homologs like LlaI and BsuMI are distinguished by a unique three-gene organization and the ability to target DNA site-specifically. Here, we report the atomic-resolution crystal structures of the DNA-binding module LlaI.R1 and AAA+ motor LlaI.R2 from the Lactococcus lactis LlaI anti-phage defense system. The crystallized LlaI.R2 hexamer traps two distinct active site conformations that correlate to different states of the nucleotide hydrolysis cycle and reveal that the organization of the critical catalytic machinery present in canonical McrB homologs is also conserved in non-canonical R2 proteins. Although canonical McrB homologs are strictly GTP-specific, we find that the R2 proteins from LlaI and BsuMI do not discriminate between different nucleotides, even when in complex with their respective R1 partners. Using mutagenesis, we define surfaces on the LlaI.R1 structure that are critical for DNA-binding and interaction with LlaI.R2. These observations support computational modelling of the assembled LlaI restriction system bound to DNA. Together, our data provide new insights into the evolution of nucleotide specificity in McrBC restriction complexes and the molecular mechanisms governing McrBC-catalyzed DNA translocation and cleavage.

15
The first chromosome-scale genome assembly of Blumeria graminis f. sp. avenae provides insights into genome evolution and host specialization

Ding, Y.; Zhang, P.; Ociepa, T.; Nucia, A.; Guan, H.; Kowalczyk, K.; Park, R. F.; Okon, S.

2026-08-30 genomics 10.64898/2026.08.28.747853 medRxiv
Top 3%
1.6%
Show abstract

Blumeria graminis f. sp. avenae (Bga), the causal agent of oat powdery mildew, is one of the most host-specialized members of the B. graminis species complex. Despite its agricultural importance, the lack of a high-quality reference genome has limited studies of host specialization, virulence evolution and comparative genomics in this pathogen. Here, we generated the first chromosome-scale genome assembly of Bga using an integrative approach combining long- and short-read sequencing, Hi-C scaffolding and transcriptome data. The Bga genome exhibits hallmark features of powdery mildew fungi, including extensive repeat content and low gene density. Comparative analyses revealed that genome expansion is primarily associated with historical transposable element proliferation rather than recent transpositional activity. Genome organization is consistent with a functionally stratified "one-speed" model, in which genes associated with pathogenicity, including predicted effectors and infection-responsive genes, are preferentially located in transposable element-rich regions characterized by reduced synteny conservation and extended intergenic spaces. In contrast, conserved genes are concentrated in compact genomic regions and maintain strong syntenic conservation across cereal-infecting formae speciales. Hi-C analyses demonstrated a highly structured chromatin architecture and revealed genome organization patterns associated with infection-related gene expression. Comparative genomic analyses indicated that host specialization in Bga is driven by localized diversification of a relatively small subset of genes rather than large-scale genome restructuring. These results provide the first high-quality genomic resource for Bga and offer new insights into the evolutionary mechanisms underlying host specialization in powdery mildew fungi.

16
Bayesian adaptive experimental design for efficient microbial genome-wide association studies

Helekal, D.; Blomqvist, S. O. P.; Mukherjee, A.; Bowcutt, B. A.; Palace, S. G.; Grad, Y. H.

2026-08-31 genetics 10.64898/2026.08.26.747358 medRxiv
Top 4%
1.4%
Show abstract

Bacterial genome-wide association studies (GWAS) offer a powerful approach to identify the genetic basis of a trait measured in a set of sequenced isolates. As the number of sequenced isolates has grown, the limiting factor for GWAS has become phenotyping enough isolates to achieve statistical power. To overcome the need for large-scale phenotyping, we developed Bayesian Adaptive Sequential Sampling GWAS (BASS-GWAS), which couples Bayesian adaptive experimental design with a sparse regression model to select maximally informative isolates for phenotypic testing. BASS-GWAS efficiently recovered causal loci for three antimicrobial resistance traits in Neisseria gonorrhoeae, requiring many fewer phenotyped isolates than random sampling. We applied BASS-GWAS to discover variants enabling gyrBD429N-dependent cross-resistance to the novel topoisomerase inhibitors zoliflodacin and gepotidacin. After phenotyping fewer than 30 isolates, we identified and then validated both parCD86N and a gyrA-parE-based pathway as enabling cross-resistance. BASS-GWAS provides a practical and statistically principled solution for efficient bacterial GWAS.

17
Temporal, genome-scale analysis of Myxococcus xanthus developmental fate in a mixed population

Mittal, S.; Mandal, S.; Farrugia, M. A.; Crosson, S.; Fiebig, A.; Kroos, L.

2026-08-31 molecular biology 10.64898/2026.08.28.747804 medRxiv
Top 4%
1.4%
Show abstract

Myxococcus xanthus bacteria form aggregates when starved on solid surfaces and some cells differentiate into spores. Studies of mutants in monoculture have advanced knowledge of this multi-cellular developmental process, but our understanding of the genetic determinants is incomplete. To assess gene function genomewide, we generated a pool of barcoded transposon insertion mutants, subjected it to starvation, and separated developmental samples into non-aggregated cells, aggregated cells, and spores. We also subjected our pool to chemically-induced unicellular sporulation. Evaluation of changes in the abundance of mutants in samples allowed identification of 200 genes in which insertions reproducibly caused distinct patterns of depletion and/or accumulation over time. Many of these genes have well-established roles in development, validating our approach, while many others have not previously been associated with development. Genes involved in type IV pili (T4P)-dependent motility were more important than gliding motility genes for aggregation and sporulation in the mixed population. Although exopolysaccharide (EPS) synthesis genes are required for aggregation in monoculture, most were dispensable for aggregation in our pool, consistent with EPS sharing between cells, yet these genes were required cell-autonomously for efficient sporulation. Genes for positive regulators of EPS synthesis were important for aggregation as well as sporulation, suggesting functions beyond EPS production. Insertions in several novel genes impaired both starvation- and chemically-induced sporulation. Many genes increased the efficiency of starvation-induced sporulation. Some of these mutants, which we call "developmental winners", are novel cheaters. Our results demonstrate the power of using the newly-created mutant library to elucidate M. xanthus biology.

18
The RNA virome of early metazoans sheds light on long-term virus-host relationships

Ortiz-Baez, A. S.; Mifsud, J. C. O.; Schwarz, J.; Sadiq, S.; Holmes, E. C.

2026-08-31 microbiology 10.64898/2026.08.30.748148 medRxiv
Top 4%
1.3%
Show abstract

Ctenophores and placozoans arose early in metazoan evolution and are characterized by traits associated with key aspects of animal evolution. Despite the evolutionary significance of ctenophores and placozoans, their RNA viromes are poorly understood. To determine the diversity and evolution of RNA virome in these organisms, particularly whether the viruses present with these ancient host lineages similarly occupy basal phylogenetic positions, we analysed publicly available transcriptome data from the Sequence Read Archive (SRA). Accordingly, we identified 26 putative novel viruses classified into 11 virus groups, including members of the families Flaviviridae and Chuviridae. The novel viruses clustered with those previously identified in vertebrates, invertebrates, plants and fungi. Notably, some virus sequences within the Flaviviridae, Chuviridae, Lispiviridae and Marnaviridae were highly divergent, branching deeply relative to their closest known relatives or forming distinct lineages, in some cases suggesting a divergence early in metazoan evolution. In contrast, viruses within the Birnaviridae, Endornaviridae, Mymonaviridae, Narnaviridae, Phasmaviridae, Orthomyxoviridae, Orthototiviridae, and some viruses within the Picornavirales, exhibited patterns consistent with more recent diversification and host jumping. In addition, RNA viruses were detected across multiple species and tissues within the Ctenophora (including whole organisms and embryos) and Placozoa, expanding their host range and highlighting a largely uncharacterized diversity. Together, these findings expand the known diversity and host range of several virus groups, and shed light on virus evolution in early metazoans, demonstrating both host jumping within aquatic environments and virus host-associations that may span the entirety of animal evolution.

19
Symbiont spatial organisation is dynamically regulated within cnidarian host tissues

Jilani, A.; Allgeyer, E. S.; Li, X.; Guo, M.; Sevilgen, D. S.; Ball, A.; Xiong, F.; McLaren, S. B. P.

2026-08-31 developmental biology 10.64898/2026.08.28.743919 medRxiv
Top 4%
1.1%
Show abstract

The symbiosis with photosynthetic dinoflagellate algae enables corals to build and sustain reef ecosystems. Individual coral polyps hold algal symbionts in their epithelial endoderm cells and lose them under environmental stress, leading to coral bleaching. How the host integrates symbionts into its body plan is not well understood. Here, using a combination of high-resolution imaging, quantitative analysis, and environmental perturbations in the sea anemone Exaiptasia diaphana (Aiptasia) and reef-building coral Pocillopora damicornis, we uncover a spatial organisation of symbionts along the aboral-oral axis of cnidarian polyps that emerges under the long-range translocation of symbionts between host cells through a fluid-filled cavity. The symbiont distribution becomes specifically enriched in the tentacle bud endoderm during Aiptasia polyp morphogenesis. This pattern can form in darkness and with algae-sized inert spheres, suggesting an innate host-intrinsic mechanism. Symbiont-occupied host cells are mechanically constrained within the endoderm and thus unable to rearrange; instead, they go through cycles of symbiont expulsion and re-uptake via the host gastric cavity, with regionally biased rates of these behaviours providing a route to enrich symbionts in the tentacles. Symbiont organisation is remodelled under increased light in adult coral polyps, with a characteristic pattern of reduced tentacle enrichment, lateral clustering and retention in the body column emerging over a timescale of days. Together, our findings reveal that the spatial organisation of symbionts is dynamically regulated in cnidarian host tissues, a capacity that may shape both the establishment of symbiosis and its resilience under environmental change.

20
Loss of replication and transcription systems accompanying transition to nucleus-dependent replication in Ariadnavirales, a proposed new order in nucleocytoviricot class Megaviricetes

Yutin, N.; Wolf, Y. I.; Krupovic, M.; Koonin, E. V.

2026-08-30 evolutionary biology 10.64898/2026.08.29.747986 medRxiv
Top 4%
1.1%
Show abstract

Sicyoidochytrium minutum DNA virus (SmDNAV) was isolated several years ago from a protist host of family Thraustochytriaceae of the class Labyrinthulomycetes. This virus shared little similarity to other viruses in gene content and protein sequences, albeit seemingly belonging to the phylum Nucleocytoviricota. By extensive searches in genomic and metagenomic sequence databases, we identified numerous long contigs related to the SmDNAV genome and analyzed proteins shared by these putative viruses. Phylogenetic analyses place these viruses within the class Megaviricetes, outside of all established orders, and as a sister group to the clade combining families Mamonoviridae and Manesviridae. Homologs of SmDNAV proteins were found in association (either integrated or co-sequenced) with other Labyrinthulomycetes and Rhodophyta protists from diverse marine and freshwater environments. Consequently, we propose SmDNAV as the prototype member of a new order, provisionally named Ariadnavirales, within class Megaviricetes, phylum Nucleocytoviricota. Members of Ariadnavirales have lost most of the genes encoding components of the replication and transcription systems that are otherwise conserved in nucleocytoviricots, suggestive of transition to genome replication and expression dependent on the host nucleus.